# Sprint plan: the 12-week demo Oct 1, 2026 · six two-week sprints, 5 Oct to 25 Dec 2026 The demo is ready for dry runs on 18 December and demo-ready on 23 December. Committed engineering work is about 35 engineer-weeks (173.5 engineer-days) against 48 available (72%). The rest is buffer for real-data surprises. That buffer exists only because the thin golden path is already built; the original plan had none. **Inputs:** - `BACKLOG.md` - `docs/PRD.md` §6, as revised on 1 Oct - `docs/eng-review.md`, tasks E-T1 to E-T12 - `docs/design-review.md`, tasks D-T1 to D-T9 - `docs/build-plan.md`, Results Review tasks are recommendations until the owner accepts them. Sprint 1 planning confirms them; the ones that block sprint 1 are listed under "Decisions needed on 5 October". ## Team and capacity | Role | Short | Engineering capacity | Main lane | |---|---|---|---| | Data engineer | DE | 10 days a sprint | Hospital zone: intake, privacy, export, server | | Machine-learning engineer 1 | ML1 | 10 days | Documents, gateway, extraction, digest, synopsis | | Machine-learning engineer 2 | ML2 | 10 days | Estimator, replay, forecast, simulator, calibration | | Full-stack engineer | FS | 10 days | Web app in both zones, screens, export formats | | Product and clinical lead | PL | not counted | Study access, gold labels, buyers, demo script | | Biostatistician | BS | not counted | Methods, hand extractions, report content, dry run 1 | | Regulatory specialist, part time | RA | not counted | Permissions, consent scope, theme labels | | Clinician from the study's field | CL | not counted | Review queue, elicitation, dry run 2 | A sprint offers 40 engineer-days. One day per person goes to planning, review and retro, which leaves 36. The plan commits about 30 a sprint (about 75% of 40) and names stretch items to pull if a sprint runs ahead. **There is no designer on the team.** Design tasks go to the product lead, working in the design canvas, and to FS. This is a risk, listed below. ## Overview | Sprint | Dates | Goal | Committed (eng-days) | Gate at the end | |---|---|---|---|---| | 1 | 5–16 Oct | Foundations, privacy hardening, and the go or no-go on real data | 30.5 | **Go or no-go on real data (16 Oct)** | | 2 | 19–30 Oct | Real documents in; the web app takes its shape | 32 | Model extraction runs on surrogate Word and PDF files | | 3 | 2–13 Nov | Acts 1 and 2 on the real study | 33 | F1.2 and F1.3 accuracy measured against gold sets | | 4 | 16–27 Nov | Act 3: the real profile, reviewed by clinicians | 27 | First approved real export; ≥ 80% of parameters reviewed | | 5 | 30 Nov–11 Dec | Act 4: from a brief to a signed report | 28 | Signed report exports in Word and PDF | | 6 | 14–25 Dec | Act 5, two dry runs, and fixes | 23 | Demo-ready on 23 Dec | ``` Oct 5 Oct 19 Nov 2 Nov 16 Nov 30 Dec 14 Dec 25 |---- S1 ----|---- S2 -----|---- S3 -----|---- S4 -----|---- S5 -----|--- S6 ----| privacy, documents, acts 1-2 act 3 act 4 act 5, spikes, web shell on real data profile + brief to dry runs GO/NO-GO ▲ review ▲ report ▲ demo ▲ SURPASS summary due Nov 30 (PL, time-boxed) ``` ## Sprint 1 · 5–16 October **Sprint goal:** close the three privacy gaps that block real data, measure extraction on surrogate Vietnamese documents, and reach a go or no-go on real data by 16 October. | Priority | Item | Estimate | Owner | Dependencies | |---|---|---|---|---| | P0 | E-T1 git init, first commit, CI (pytest, 30-seed calibration, golden path) | 1d | DE | none | | P0 | E-T2 sign the export and its approval; the platform verifies | 3d | DE | E-T1 | | P0 | E-T3 quarterly series and the differencing check in the gate | 3d | ML2 | E-T1 | | P0 | E-T4 single-site rule and per-site consent flag | 3d | ML2 | E-T1 | | P0 | Backlog 0.4: extraction spike on at least 5 surrogate Vietnamese protocols, scans measured for F1.10 | 3d | ML1 | surrogate documents from PL | | P0 | Backlog 0.5: map a raw dataset to the clinical schema, 50 variables, time recorded | 2d | DE | a de-identified surrogate dataset | | P0 | Front-end framework chosen; app skeleton with the four roles | 5d | FS | decision on 5 Oct | | P1 | Word and PDF text extraction with page and paragraph anchors (first half) | 5d | ML1 | none | | P1 | E-T9 NFC normalisation at intake | 1d | DE | none | | P1 | Elicitation form for between-study spread and gap fills, with BS | 2d | ML2 | BS | | P1 | D-T4 DESIGN.md: scales and component inventory | 2d | FS | none | | P2 | E-T12 `ClinicalStore.close()`; drop the job queue from the stack | 0.5d | FS | none | | Stretch | E-T5 allow-list de-identification, started | (5d) | DE | none | **Load:** 30.5 days committed (76% of 40). DE 7, ML1 8, ML2 8, FS 7.5. **Non-engineering:** - PL and RA run the study inventory and permission check (backlog 0.7). PL files the hospital server request in week 1 (backlog 0.8). - PL starts the ten price interviews (PRD §8). - BS pre-registers the replay quantities and the calibration seeds before any real data is seen. **Sprint-specific risk:** the permission check runs past 16 October. Mitigation: keep sprint 2 data-agnostic and name the partner trial fallback (PRD §6) by 9 October. ### Sprint 1 result (run 1 Oct) All engineering items are done except the two spikes, which are blocked on surrogate documents and a surrogate dataset that PL supplies. The stretch item E-T5 was pulled in. Three items came forward from later sprints in thin form: the act-grouped rail, the clinician review queue, and the sign-off step. CI is green with 100 tests. The board is live at `/sprint` in the demo app (`python -m demo_app`). The gate, go or no-go on real data, is still open and is not an engineering decision. ## Sprint 2 · 19–30 October **Sprint goal:** a real Word or PDF protocol goes in and comes out as structured fields with page-level sources, and the web app shows acts 1 and 2. | Priority | Item | Estimate | Owner | Dependencies | |---|---|---|---|---| | P0 | E-T5 allow-list de-identification with the golden-path regression test | 5d | DE | E-T1 | | P0 | Hosted model behind the gateway; prompts and schemas in version control | 3d | ML1 | provider decision | | P0 | Model field extraction v1, compared with the rule baseline | 5d | ML1 | text with anchors | | P0 | Port the package, retrospective and feedback screens to the framework | 5d | FS | skeleton | | P1 | E-T7 two command-line entry points; golden path as two processes | 3d | DE | E-T2 | | P1 | E-T8 site types read from the data | 3d | ML2 | E-T4 | | P1 | Estimator hardening for real data: visit windows from the study record, missing dates | 4d | ML2 | none | | P1 | D-T2 five-act rail and the hospital-zone UI skeleton | 3d | FS | skeleton | | P2 | E-T11 observed effect gated by data-use terms | 1d | ML2 | E-T2 | | Stretch | Real dataset mapping, if the gate said go | (4d) | DE | go on 16 Oct | **Load:** 32 days (80%). **Non-engineering:** PL, RA and CL label the gold sets (100 files for typing, 200 protocol fields), and PL schedules the investigator debriefs. ### Sprint 2 result (run 1 Oct) The goal is met on synthetic protocols: Word, PDF and Markdown, in English and Vietnamese, come out as structured fields that link to their paragraph or PDF line. The hosted model and model extraction are built and tested against a fake client. Neither has been called live, because there is no API key here, and neither is measured, because there is no gold set yet. Both are marked partial. Done: E-T7 (two processes), E-T8 (site types from data, which found and fixed a forecast bug), estimator hardening, E-T11, and the hospital-zone console (pulled forward from sprint 3). The package page was redesigned (design #4). The owner added two items mid-sprint, and both are done: - Synthetic studies for three diseases with three endpoint kinds (continuous, binary, time to event). The simulator and report handle all three. - A protocol tool that starts from an indication or from a published trial (FLAURA, KEYNOTE-024, KEYNOTE-189, CHANCE). Measured: - Calibration on 150 synthetic studies: parameter intervals 90.3% against a stated 90%, replay 86.0% against 80%, back-test 75.3% against 80%. The back-test is slightly overconfident, and that is recorded, not tuned. - 147 tests pass, and CI is green. Board: `/sprint/2` in the demo app; data in `docs/sprints/sprint-2.json`. ## Sprint 3 · 2–13 November **Sprint goal:** acts 1 and 2 run on the real study, and F1.2 and F1.3 are measured against hand-labelled sets. | Priority | Item | Estimate | Owner | Dependencies | |---|---|---|---|---| | P0 | Real dataset mapped to the clinical schema (F1.2 datasets, 50-variable gold set) | 6d | DE | go; E-T5 | | P0 | E-T10 hospital server live: encrypted disk, key outside the data folder, startup check | 3d | DE | server approval | | P0 | Extraction measured on the 200-field gold set; low-confidence field review API | 5d | ML1 | gold labels | | P0 | Retrospective on the real study (F1.3), checked against BS's 30 items | 3d | ML2 | real export | | P0 | De-identification review and export approval screens in the hospital zone | 4d | FS | D-T2 | | P1 | File typing on real formats; 100-file gold harness (F1.1) | 3d | ML1 | gold labels | | P1 | Amendments and review timelines from real documents | 2d | ML2 | extraction v1 | | P1 | Low-confidence field review screen | 3d | FS | review API | | P1 | D-T3 states for acts 1 and 2: empty, error, partial | 2d | FS | DESIGN.md | | P2 | Harness for the 30 retrospective items | 2d | ML2 | BS's hand extraction | **Load:** 33 days (83%). This is the heaviest sprint; its stretch is empty on purpose. **Non-engineering:** BS hand-extracts the 30 retrospective items, and RA labels themes on the real letters. **If the gate said no-go:** acts 1 and 2 run on the partner trial or on synthetic data. DE's 6 days go to a second synthetic package with three site types and Word and PDF documents. ## Sprint 4 · 16–27 November **Sprint goal:** the real profile crosses the export gate, and investigators review at least 80% of it. | Priority | Item | Estimate | Owner | Dependencies | |---|---|---|---|---| | P0 | First real export: steward approval, signature, import, and fixes | 3d | DE | E-T2, E-T3, E-T4 | | P0 | Clinician review queue: confirm, correct, reject, keyboard keys (F1.6, D-T6) | 5d | FS | profile API | | P0 | Profile on the real export, plus the elicitation session (between-study spread, gap fills) | 3d | ML2 | CL session | | P0 | Forecast by site type on the real study, with back-test (F5.4) | 3d | ML2 | real export | | P1 | Profile screen: within-study interval beside the prediction interval | 3d | FS | review queue | | P1 | Feedback digest: model themes behind the gateway, human review, machine-translated Vietnamese quotes (F1.4) | 6d | ML1 | gateway | | P1 | Audit log screen | 2d | DE | none | | P2 | Calibration re-run at 150 seeds; trust panel updated | 1d | ML2 | none | | P2 | Theme gold harness | 1d | ML1 | RA labels | | Stretch | Real-data quality fixes | (3d) | DE | none | **Load:** 27 days, plus the 3-day stretch held for real-data fixes (68–75%). **Non-engineering:** - CL and two investigators run the elicitation and the review queue. - PL writes the SURPASS solution summary, due 30 November, time-boxed to two days (PRD §9). ## Sprint 5 · 30 November–11 December **Sprint goal:** a new brief becomes a synopsis, is simulated twice, and ends in a signed report a sponsor could buy. | Priority | Item | Estimate | Owner | Dependencies | |---|---|---|---|---| | P0 | Brief to synopsis as USDM data, each element traced (F2.1) | 4d | ML1 | none | | P0 | Simulator driven by the synopsis: visit schedule, local activation if the documents show it | 4d | ML2 | F2.1 | | P0 | D-T1 simulate screen with the "What changes" column and the forecast chart | 3d | FS | simulator API | | P0 | Design report in Word and PDF (F2.8) | 4d | DE | report model | | P0 | Brief screen (step 7) | 3d | FS | F2.1 | | P1 | D-T9 sign-off as a closing step; export unlocks after the signature | 2d | DE | Word and PDF export | | P1 | Report content review with BS | 3d | ML2 | BS | | P1 | D-T5 presentation mode | 2d | FS | DESIGN.md | | P2 | Synopsis text drafted through the gateway with source links | 3d | ML1 | gateway | **Load:** 28 days (70%). **Non-engineering:** BS reviews the report content, and CL reviews the synopsis. ## Sprint 6 · 14–25 December **Sprint goal:** replay and calibration on the real study, two dry runs, every unexplained number fixed, demo-ready on 23 December. | Priority | Item | Estimate | Owner | Dependencies | |---|---|---|---|---| | P0 | Replay on the real study with the pre-registered quantities, plus the calibration panel (F3.8) | 3d | ML2 | real data | | P0 | Dry run 1 with BS (17 Dec): every unexplained number logged and fixed | 6d | all | all acts | | P0 | Dry run 2 with an investigator (21 Dec): defects logged and fixed | 6d | all | dry run 1 | | P0 | Golden path green in CI on synthetic data, and on the real study inside the hospital zone | 2d | DE | E-T7 | | P1 | Limits screen and the remaining states | 3d | FS | none | | P1 | Extraction fixes from the dry runs | 3d | ML1 | none | | Stretch | Design critique (F2.5), the first cut | (5d) | ML1 | none | | Stretch | Evidence file (F9.3) | (4d) | ML2 | none | | Stretch | "Ways to recover the timeline" levers | (3d) | FS | owner decision | **Load:** 23 days (58%). This sprint is light on purpose: dry-run fixes always take more than planned. 24 and 25 December are treated as half days. ## Cut order if time runs short The order follows PRD §6: 1. The design critique: it is not in the committed plan. 2. Model-based theming: keep the keyword baseline with human review. 3. The Vietnamese interface beyond the report itself. Replay and calibration are protected, and so are the three privacy tasks (E-T2, E-T3, E-T4) and E-T5. ## Risks | Risk | Impact | Mitigation | |---|---|---| | No study qualifies by 16 Oct | Acts 1–3 lose real data | Partner trial fallback named by 9 Oct; a second synthetic package with Word and PDF documents in sprint 3 | | The hospital server is not approved in time | No real data on our hardware | Run in the hospital's own environment under the steward, with the same encrypted-disk checks; start the request on 5 Oct | | Extraction on real protocols is below 90% | F1.2 acceptance missed | Report the measured number (CLAUDE.md); review every field by hand for the demo | | No designer on the team | Undrawn screens are built as plain forms | PL draws them in the design canvas in sprints 1–3; DESIGN.md first | | One full-stack engineer owns every screen in both zones | A sick week stalls acts 3–5 | Keep screens on DESIGN.md components; ML1 can take the brief screen | | The clinician's time | F1.6 below 80% reviewed | Book the review and elicitation sessions in sprint 2 for sprint 4 | | The SURPASS summary pulls the lead in late November | Gold labels and the elicitation slip | Two-day time-box (PRD §9); BS covers the elicitation | | Sending sponsor documents to a hosted model | A confidentiality breach without patient data | Gateway use for sponsored documents only with written permission; IIT documents first | | The Christmas week | Dry-run fixes slip | Demo-ready target is 23 Dec; 24 and 25 Dec are buffer | ## Definition of done Applies to every item in every sprint: - [ ] The item names its PRD feature ID or review task ID. - [ ] Its acceptance test is written with it. If a threshold cannot be met, the measured number is reported and the test is not changed. - [ ] The golden path stays green in CI; calibration does not fall below its last recorded coverage. - [ ] Every number the item shows carries a label, an interval and provenance. - [ ] Text a user sees exists in Vietnamese and English. - [ ] No real patient data in the repository, logs or fixtures. - [ ] Code reviewed and merged. - [ ] Docs updated where behaviour changed. ## Key dates | Date | Event | |---|---| | Mon 5 Oct | Sprint 1 planning: confirm review tasks, choose the framework, file the server request | | Fri 9 Oct | Partner trial fallback named | | Fri 16 Oct | **Go or no-go on real data**; sprint 1 review and retro | | Fri 30 Oct | Sprint 2 review: extraction on Word and PDF | | Fri 13 Nov | Sprint 3 review: F1.2 and F1.3 measured | | Fri 27 Nov | Sprint 4 review: first approved real export, review coverage | | Mon 30 Nov | SURPASS solution summary due | | Fri 11 Dec | Sprint 5 review: signed report | | Thu 17 Dec | Dry run 1, biostatistician | | Mon 21 Dec | Dry run 2, investigator | | Wed 23 Dec | Demo-ready | | Fri 25 Dec | Sprint 6 retro; demos start in January | Mid-sprint check-ins fall on the second Monday of each sprint: 12 Oct, 26 Oct, 9 Nov, 23 Nov, 7 Dec and 21 Dec. ## Decisions needed on 5 October These block sprint 1. Each has a recommendation in the review named. | Decision | Recommendation | Source | |---|---|---| | Front-end framework | One the full-stack engineer already knows well; the static pages port to any of them | CLAUDE.md, open decisions | | Sign the export (A1), stop differencing (A2), the single-site rule (A3) | Yes to all three | eng-review.md | | Allow-list de-identification (C1) | Yes | eng-review.md | | Git and CI (C6) | Yes, today | eng-review.md | | Hosted model provider for documents without patient data | Choose in sprint 1, wire in sprint 2 | demo-architecture.md | | A hospital-zone UI separate from the platform app (design #2) | Yes | design-review.md | The other review findings are decided at the sprint that schedules them.